Search CORE

2 research outputs found

A DOM-Tree based Representation of Web Document Structure for Web Mining Applications

Author: Manoj Kumar Sarma, Anjana Kakoti Mahanta
Publication venue: 'Auricle Technologies, Pvt., Ltd.'
Publication date: 30/06/2017
Field of study

Among the three broad areas of Web mining, Web Structure Mining is the method of discovering structure information from either the web hyperlink structure or the web page structure. In order to apply data mining techniques on web pages, a good and efficient representation of web pages is required that could depict the actual hierarchical structure of web pages. The work presented here aims to find out a representation of web documents that could be used as input for different data mining techniques. The present research work further aims at applying this representation for efficient clustering of web documents where clustering will be performed based on not only the web page content but also the structural layout of a web page

International Journal on Recent and Innovation Trends in Computing and Communication

Study on Distance Measures for Clustering of Web Documents based on DOM-Tree based Representation of Web Document Structure

Author: Manoj Kumar Sarma, Anjana Kakoti Mahanta
Publication venue: 'Auricle Technologies, Pvt., Ltd.'
Publication date: 30/06/2017
Field of study

Among the three broad areas of Web mining, Web Structure Mining is the method of discovering structure information from either the web hyperlink structure or the web page structure. In order to apply data mining techniques on web pages, a good and efficient representation of web pages is required that could depict the actual hierarchical structure of web pages. The work presented here aims to find out an appropriate distance measure (also called as similarity measure) for strings that can be used for clustering of web documents and also for other data mining applications

International Journal on Recent and Innovation Trends in Computing and Communication